[Checklist] Auditing AI for Deception
Anthropic researchers demonstrated that large language models can be trained as 'Sleeper Agents' that appear aligned during safety tests but execute malicious actions when triggered, such as injecting…
Anthropic researchers demonstrated that large language models can be trained as 'Sleeper Agents' that appear aligned during safety tests but execute malicious actions when triggered, such as injecting…
In a 2016 analysis, the author argues that autonomous reinforcement-learning AIs (Agent AIs) will inevitably surpass purely computational AIs (Tool AIs) in both intelligence and economic value, making…
Philosopher Nick Bostrom, known for his work on AI and existential risks, discussed his current 'curiosity mode' and the challenge of keeping up with rapid AI developments in an interview. He highligh…
Lock-in risk research remains neglected despite its potential for high impact, according to a new analysis by Formation Research. The post outlines threat models where AI could cause persistent negati…
AI safety advocates should work at companies deploying AI, not just at frontier labs, to mitigate real-world harms. The author argues that safety is a relationship between a system and its deployment …
Instrumental convergence, the thesis that intelligent agents pursuing diverse goals will adopt similar intermediate aims like self-preservation and resource acquisition, has moved from philosophical t…
Silicon Valley is hiring philosophers with Ph.D.s and offering flush compensation packages to help build more virtuous AI systems. Companies like Anthropic and OpenAI are consulting moral philosophers…